【文章标题】:[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law
【文章标题】:[AINews] 参数的消亡:智谱AI CEO唐杰谈GLM 5.3与新的后训练扩展定律
【文章正文】: We’ve covered
GLM 5.2
very excitedly before, and Prof Jie Tang’s belief that there will be an
open weights Fable-class model by end of the year
(
spot check - with 134 days left, there are now two 2-3T models (
Qwen 3.8 Max
and
Kimi K3
) with estimates that Fable is
3-7T
, and only
2 points higher on the AA index
.)
我们之前非常兴奋地报道过
GLM 5.2
,以及唐杰教授相信年底前会出现
开源权重的Fable级模型
(
抽查一下——距离年底还有134天,现在已有两个2-3T模型(
Qwen 3.8 Max
和
Kimi K3
),估计Fable有
3-7T
,在AA指数上仅高出
2分
。)
Prof Jie Tang is back on X
to tell us that our shorthand for model sizes is no longer enough: ”
Parameter count is only meaningful alongside three others — how much
data
you have, where you intend to spend your
compute
, and who will run the model,
under what conditions
.”
唐杰教授回到X平台
告诉我们,我们用来表示模型大小的简写已经不够了:“
参数数量只有在与其他三个因素一起考虑时才有意义——你拥有多少
数据
,你打算将
算力
花在哪里,以及由谁来运行模型,
在什么条件下
运行。”
We have covered
Chinchilla
(and
post-Chinchilla
) scaling laws in past LS years, but, so we will skip the history lesson, but it is good to level-set on why Chinchilla’s assumptions were wrong in the
Inference Inflection
world (no fixed number, between 200-900 toks/param, citing
Roberts et al
on task dependence).
我们在过去的LS年中已经报道过
Chinchilla
(和
后Chinchilla
)扩展定律,但这里我们就不上历史课了,不过有必要统一认识,了解为什么Chinchilla的假设在
推理拐点
世界中是错误的(没有固定数字,在200-900 tokens/参数之间,引用了
Roberts等人
关于任务依赖性的研究)。
In short: Memorization prefers more parameters. Reasoning prefers more post-training data and
effective depth
.
简而言之:记忆更偏好更多参数。推理更偏好更多后训练数据和
有效深度
。
GLM-5.3’s
big jumps come solely from RL on long horizon environments:
GLM-5.3
的巨大跃升完全来自长时程环境上的强化学习:
The environments now cover a much broader range of production workflows, with tasks designed around how engineering and research work is actually carried out in practice.
这些环境现在涵盖了更广泛的生产工作流,任务设计围绕工程和研究工作在实际中是如何开展的。
Some represent several days of work for an experienced engineer.
有些任务相当于一位经验丰富的工程师数天的工作。
In an ML infrastructure task, for example, the model may be given the same working environment as an engineer, with
access to compute clusters, storage systems, internal documentation, codebases, and experiment results
. It must diagnose bottlenecks across the training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup while preserving correctness. Training on environments at this level pushes the model toward taking
ownership of substantial work end to end
, rather than relying on users to decompose the problem and supervise each step.
例如,在ML基础设施任务中,模型可能会获得与工程师相同的工作环境,包括
访问计算集群、存储系统、内部文档、代码库和实验结果
。它必须诊断训练栈中的瓶颈、实施优化、运行实验,并在保持正确性的同时交付可衡量的端到端加速。在这种水平的环境上训练,会推动模型对
端到端的重要工作负起责任
,而不是依赖用户来分解问题并监督每一步。
For those following
the recursive self improvement story
, their entire environment and judging and verifier process is synthetic all the way down:
对于那些关注
递归自我改进
故事的人来说,他们的整个环境以及评判和验证过程完全都是合成的:
As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment. A useful task environment has to be executable, verifiable, and close to real professional work — and we need many of them, not a handful of hand-built ones. To scale this process, we built
pipelines that synthesize environments end to end
, and for a subset of tasks, the RL reward signal as well.
随着智能体能力的提升,扩展后训练的很大一部分难度从模型转移到了环境。一个有用的任务环境必须是可执行、可验证且接近真实专业工作的——而且我们需要很多这样的环境,而不是少数几个手工构建的。为了扩展这一过程,我们构建了
端到端合成环境的流水线
,并且对于一部分任务,还包括RL奖励信号。
Research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state
; a judge agent then attempts each task to verify that it is actually solvable. Verifiers are synthesized without access to the reference solution, while solver trajectories are used to discover and close reward shortcuts. A verifier that passes oracle, no-op, and unsolved-state checks produces a binary reward reliable enough to train on directly.
研究智能体从真实工作中收集任务模式,并将其转化为可运行的长时程环境,具有多步依赖和隐藏状态
;然后一个评判智能体尝试每个任务,以验证它确实可解。验证器在不访问参考解的情况下被合成,同时求解器轨迹被用来发现和关闭奖励捷径。一个通过了oracle、无操作和未解决状态检查的验证器,会产生一个足够可靠的二元奖励,可以直接用于训练。
To put an end to parameter count obsesssion, Prof Jie identifies 5 knobs of scaling, including MoE sparsity with the
new XA-YB notation
. He notes that advanced skills (e.g., finding software vulnerabilities) are not retrieval/memorization problems.
为了终结对参数数量的痴迷,唐杰教授指出了5个扩展旋钮,包括使用
新的XA-YB记号
的MoE稀疏性。他指出,高级技能(例如发现软件漏洞)不是检索/记忆问题。
They require carrying long causal chains (20+ inference steps) without losing the thread.
它们需要在20步以上的推理步骤中携带长因果链而不丢失线索。
This ability does not live in total parameter count once a certain knowledge-holding threshold is reached.
一旦达到一定的知识持有阈值,这种能力就不再取决于总参数数量。
And it looks like there is much more to go.
而且看起来还有更多路要走。
AI News for 8/18/2026-8/19/2026. We checked 12 subreddits,
544 Twitters
and no further Discords.
2026年8月18日至8月19日的AI新闻。我们检查了12个子论坛、
544个Twitter账号
,没有额外的Discord。
AINews’ website
lets you search all past issues. As a reminder,
AINews is now a section of Latent Space
. You can
opt in/out
of email frequencies!
AINews的网站
可以让你搜索所有过往期刊。提醒一下,
AINews现在是Latent Space的一个栏目
。你可以
选择接收/退订
邮件频率!
AI Twitter Recap
AI Twitter回顾
Open-Weight Models, Compression, and Benchmark Movement
开源权重模型、压缩与基准测试动态
Ornith-1.5 lands as a serious new open family
:
@ornith_
released
Ornith-1.5
in
9B dense, 35B MoE, and 397B MoE
variants under
MIT
, with quantized formats including
FP8, GGUF, MLX, and NVFP4
. The headline claim is end-to-end
self-improvement
: the model proposes tasks, generates scaffolds, and produces RL rollouts to create new training experiences. Reported evals are strong across agentic/coding workloads, including
Terminal-Bench 2.1: 86.1
,
SWE-Bench Verified: 86
,
DeepSWE: 56
,
HLE: 44.6
, and
Tool Decathlon: 71.2
. The release was quickly wired into serving stacks by
vLLM
and
Ollama
.
Ornith-1.5作为严肃的新开源系列登场
:
@ornith_
发布了
Ornith-1.5
,提供
9B稠密、35B MoE和397B MoE
三种变体,采用
MIT
许可,量化格式包括
FP8、GGUF、MLX和NVFP4
。其头条主张是端到端的
自我改进
:模型提出任务、生成脚手架,并产生RL rollout以创造新的训练体验。据报道,在智能体/编码工作负载上的评估表现强劲,包括
Terminal-Bench 2.1: 86.1
,
SWE-Bench Verified: 86
,
DeepSWE: 56
,
HLE: 44.6
,以及
Tool Decathlon: 71.2
。该发布很快被
vLLM
和
Ollama
集成到服务栈中。
Compression continues to get more aggressive without fully collapsing utility
:
@UnslothAI
and
@danielhanchen
shipped new
Qwen3.8-27B GGUFs
using
Dynamic V3
, claiming roughly
10% higher accuracy
at the same size and releasing
1-bit quants
that still retain about
77% of BF16 accuracy
while running on
8GB RAM
. Their new
Divergence-300
metric extends top-1% greedy accuracy across longer generations using unseen examples from
Terminal Bench
,
DeepSWE
, and related tasks.
压缩继续变得更加激进,同时没有完全丧失效用
:
@UnslothAI
和
@danielhanchen
发布了新的
Qwen3.8-27B GGUF
,使用
Dynamic V3
,声称在相同大小下精度大约提高
10%
,并发布了
1-bit量化
版本,在
8GB RAM
上运行时仍保持约
77%的BF16精度
。他们新的
Divergence-300
指标使用来自
Terminal Bench
、
DeepSWE
及相关任务的未见示例,在更长生成中扩展了top-1%贪婪精度。
Agent and legal eval boards continue to reshuffle
:
@arena
published a Pareto view of
Agent Arena
, where
Claude Opus 5 (High)
leads quality, but lower-cost models like
Kimi K3
,
GLM 5.2
,
Grok 4.5
, and
GPT-5.6 Luna
define much of the value frontier. Separately,
@ValsAI
reported
Grok 4.6
at
#3/49
on
Legal Research Bench
with
48.1%
,
500k context
, tool/image/file support, and relatively low pricing. For open models,
@ValsAI
also highlighted
GLM 5.3
as
#2 on Terminal Bench
,
#3 on Legal Bench
, and
#6 on Skills Bench
among open weights.
智能体和法律评测榜单继续洗牌
:
@arena
发布了
Agent Arena
的帕累托视图,其中
Claude Opus 5 (High)
在质量上领先,但像
Kimi K3
、
GLM 5.2
、
Grok 4.5
和
GPT-5.6 Luna
这样的低成本模型定义了大量价值前沿。另外,
@ValsAI
报道
Grok 4.6
在
Legal Research Bench
上位列
#3/49
,准确率
48.1%
,支持
500k上下文
、工具/图像/文件,且定价相对较低。对于开源模型,
@ValsAI
还强调
GLM 5.3
在开源权重中排名
Terminal Bench #2
、
Legal Bench #3
和
Skills Bench #6
。
Agent Harnesses Become the New Competitive Layer
智能体框架成为新的竞争层
DeepSeek Harness’s minimalism is deliberate, not incomplete
: A detailed writeup amplified by
@ZhihuFrontier
and summarized by
@TheTuringPost
frames
DeepSeek Harness (DSH)
as an intentionally thin shell over a plugin architecture called
Cordis
. The key design choice is that
everything is a plugin
, including the agent loop itself. Early beta users reportedly shipped
100+ plugins
and filed
400+ issues
in under a week; examples range from a
gomoku model testbed
to a
database agent
that closes the SQL feedback loop by connecting the model to live query execution. The strongest takeaway is architectural: DSH is less “productized assistant” than
open agent runtime
, optimized for user-extensible tooling, swappable control loops, and business-rule injection.
DeepSeek Harness的极简主义是刻意为之,而非不完整
:由
@ZhihuFrontier
放大、@TheTuringPost
总结的一篇详细文章,将
DeepSeek Harness (DSH)
描述为一个刻意设计得很薄的壳,内部是一个名为
Cordis
的插件架构。关键设计选择是
一切都是插件
,包括智能体循环本身。据报道,早期beta用户在一周内发布了
100多个插件
并提交了
400多个问题
;例子从
五子棋模型测试平台
到
数据库智能体
,后者通过将模型连接到实时查询执行来闭合SQL反馈循环。最有力的结论是架构性的:DSH与其说是“产品化助手”,不如说是
开放智能体运行时
,针对用户可扩展工具、可替换控制循环和业务规则注入进行了优化。
TrueFoundry open-sources TrueForge and makes the harness-cost argument explicit
:
@truefoundry
,
@omarsar0
, and
@kimmonismus
all covered the launch of
TrueForge
, an
MIT-licensed
, self-hostable, vendor-neutral harness for production agents. The stack includes tool orchestration, context management, subagents, code sandboxes, human approvals, and traces, with both
local
and
hosted
deployment modes. The technical claim that resonated: on a
14-task enterprise benchmark
, TrueForge matched
Claude Managed Agents
on
Opus 4.8
while using about
30% fewer tokens
, and routing to
GLM-5.2
cut cost by around
75%
while preserving accuracy. The broader industry theme—also echoed by
@bradenjhancock
and
@dbreunig via @rseroter
—is that the
session/environment/memory/tools layer
is becoming a major source of both differentiation and savings.
TrueFoundry开源TrueForge,并明确阐述了框架成本论点
:
@truefoundry
、
@omarsar0
和
@kimmonismus
都报道了
TrueForge
的发布,这是一个
MIT许可
、可自托管、供应商中立的面向生产智能体的框架。该技术栈包括工具编排、上下文管理、子智能体、代码沙箱、人工审批和追踪,支持
本地
和
托管
两种部署模式。引起共鸣的技术主张是:在一个
14任务企业基准
上,TrueForge在
Opus 4.8
上匹配了
Claude Managed Agents
,同时使用的token大约减少
30%
,而路由到
GLM-5.2
则在保持准确性的同时将成本降低了约
75%
。更广泛的行业主题——@bradenjhancock
和
@dbreunig via @rseroter
也呼应了这一点——是
会话/环境/内存/工具层
正成为差异化和节省成本的主要来源。
Managed harnesses are also getting sharper observability and controls
:
@ClaudeDevs
added
memory support for self-hosted sandboxes
,
domain allow/block controls
for web tools, and a redesigned
multi-agent session viewer
with
minimap
,
grouped transcript
, and
cost-per-thread/session
. OpenAI, meanwhile, continues pushing the opposite angle: give teams the harness primitives to embed into their own products.
@OpenAIDevs
highlighted the
open-source Codex harness
as the runtime beneath internal tools, ops dashboards, and custom apps, while
@cursor_ai
shipped cloud-agent UX improvements around persistent goals and long-lived sessions.
托管框架也在获得更敏锐的可观测性和控制能力
:
@ClaudeDevs
增加了
自托管沙箱的内存支持
、Web工具的
域名允许/阻止控制
,以及重新设计的
多智能体会话查看器
,具备
小地图
、
分组记录
和
每线程/会话成本
。与此同时,OpenAI继续推动相反的角度:为团队提供可嵌入到自有产品中的框架原语。
@OpenAIDevs
强调了
开源Codex框架
作为内部工具、运维仪表盘和自定义应用之下的运行时,而
@cursor_ai
则围绕持久目标和长期会话发布了云智能体UX改进。
Post-Training, Mid-Training, and RL Systems Work
后训练、中训练和RL系统工作
More evidence that scaling is shifting from parameters toward training recipe quality
:
@kimmonismus
surfaced a notable claim from the
zAI/GLM
founder: progress is still scaling, but too much discourse has fixated on parameter count rather than
data quality, inference compute, and post-training
. The cited example is
GLM-5.3
, reportedly based on the same core base model/architecture as
GLM-5.2
, but improved substantially via about
one month of extra RL
.
更多证据表明,扩展正在从参数转向训练配方质量
:
@kimmonismus
提出了来自
zAI/GLM
创始人的一个显著主张:进展仍在扩展,但太多讨论固定在参数数量上,而不是
数据质量、推理算力和后训练
。所引用的例子是
GLM-5.3
,据报道它与
GLM-5.2
基于相同的核心基础模型/架构,但通过大约
一个月的额外RL
得到了大幅改进。
Microsoft’s Agent Lightning points at RL-through-the-harness as a practical recipe
:
@omarsar0
highlighted
Agent Lightning v1.0
, which connects arbitrary harnesses to RL through an endpoint proxy